agora inbox for pgsql-hackers@postgresql.org  
help / color / mirror / Atom feed
[PATCH v8] Avoid orphaned objects dependencies
399+ messages / 2 participants
[nested] [flat]

* [PATCH v8] Avoid orphaned objects dependencies
@ 2024-03-29 15:43 Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
  0 siblings, 0 replies; 399+ messages in thread

From: Bertrand Drouvot @ 2024-03-29 15:43 UTC (permalink / raw)

It's currently possible to create orphaned objects dependencies, for example:

Scenario 1:

session 1: begin; drop schema schem;
session 2: create a function in the schema schem
session 1: commit;

With the above, the function created in session 2 would be linked to a non
existing schema.

Scenario 2:

session 1: begin; create a function in the schema schem
session 2: drop schema schem;
session 1: commit;

With the above, the function created in session 1 would be linked to a non
existing schema.

To avoid those scenarios, a new lock (that conflicts with a lock taken by DROP)
has been put in place when the dependencies are being recorded. With this in
place, the drop schema in scenario 2 would be locked.

Also, after the new lock attempt, the patch checks that the object still exists:
with this in place session 2 in scenario 1 would be locked and would report an
error once session 1 committs (that would not be the case should session 1 abort
the transaction).

If the object is dropped before the new lock attempt is triggered then the patch
would also report an error (but with less details).

The patch adds a few tests for some dependency cases (that would currently produce
orphaned objects):

- schema and function (as the above scenarios)
- alter a dependency (function and schema)
- function and arg type
- function and return type
- function and function
- domain and domain
- table and type
- server and foreign data wrapper
---
 src/backend/catalog/dependency.c              |  59 ++++++++
 src/backend/catalog/objectaddress.c           |  66 +++++++++
 src/backend/catalog/pg_depend.c               |  12 ++
 src/backend/utils/errcodes.txt                |   1 +
 src/include/catalog/dependency.h              |   1 +
 src/include/catalog/objectaddress.h           |   1 +
 .../expected/test_dependencies_locks.out      | 129 ++++++++++++++++++
 src/test/isolation/isolation_schedule         |   1 +
 .../specs/test_dependencies_locks.spec        |  89 ++++++++++++
 9 files changed, 359 insertions(+)
  24.3% src/backend/catalog/
  45.3% src/test/isolation/expected/
  28.2% src/test/isolation/specs/

diff --git a/src/backend/catalog/dependency.c b/src/backend/catalog/dependency.c
index d4b5b2ade1..4271ed7c5b 100644
--- a/src/backend/catalog/dependency.c
+++ b/src/backend/catalog/dependency.c
@@ -1517,6 +1517,65 @@ AcquireDeletionLock(const ObjectAddress *object, int flags)
 	}
 }
 
+/*
+ * depLockAndCheckObject
+ *
+ * Lock the object that we are about to record a dependency on.
+ * After it's locked, verify that it hasn't been dropped while we
+ * weren't looking.  If the object has been dropped, this function
+ * does not return!
+ */
+void
+depLockAndCheckObject(const ObjectAddress *object)
+{
+	char	   *object_description;
+
+	/*
+	 * Those don't rely on LockDatabaseObject() when being dropped (see
+	 * AcquireDeletionLock()). Also it looks like they can not produce
+	 * orphaned dependent objects when being dropped.
+	 */
+	if (object->classId == RelationRelationId || object->classId == AuthMemRelationId)
+		return;
+
+	object_description = getObjectDescription(object, true);
+
+	/* assume we should lock the whole object not a sub-object */
+	LockDatabaseObject(object->classId, object->objectId, 0, AccessShareLock);
+
+	/* check if object still exists */
+	if (!ObjectByIdExist(object, false))
+	{
+		/*
+		 * It might be possible that we are creating it (for example creating
+		 * a composite type while creating a relation), so bypass the syscache
+		 * lookup and use a SnapshotSelf snapshot instead to cover this
+		 * scenario.
+		 */
+		if (!ObjectByIdExist(object, true))
+		{
+			/*
+			 * If the object has been dropped before we get a chance to get
+			 * its description, then emit a generic error message. That looks
+			 * like a good compromise over extra complexity.
+			 */
+			if (object_description)
+				ereport(ERROR,
+						(errcode(ERRCODE_DEPENDENT_OBJECTS_DOES_NOT_EXIST),
+						 errmsg("%s does not exist", object_description)));
+			else
+				ereport(ERROR,
+						(errcode(ERRCODE_DEPENDENT_OBJECTS_DOES_NOT_EXIST),
+						 errmsg("a dependent object does not exist")));
+		}
+	}
+
+	if (object_description)
+		pfree(object_description);
+
+	return;
+}
+
 /*
  * ReleaseDeletionLock - release an object deletion lock
  *
diff --git a/src/backend/catalog/objectaddress.c b/src/backend/catalog/objectaddress.c
index 7b536ac6fd..115cb1572e 100644
--- a/src/backend/catalog/objectaddress.c
+++ b/src/backend/catalog/objectaddress.c
@@ -2590,6 +2590,72 @@ get_object_namespace(const ObjectAddress *address)
 	return oid;
 }
 
+/*
+ * ObjectByIdExist
+ *
+ * Return whether the given object exists.
+ *
+ * Works for most catalogs, if no special processing is needed.
+ */
+bool
+ObjectByIdExist(const ObjectAddress *address, bool use_snapshot_self)
+{
+	HeapTuple	tuple;
+	int			cache = -1;
+	const ObjectPropertyType *property;
+
+	if (!use_snapshot_self)
+	{
+		property = get_object_property_data(address->classId);
+
+		cache = property->oid_catcache_id;
+	}
+
+	if (cache >= 0)
+	{
+		/* Fetch tuple from syscache. */
+		tuple = SearchSysCache1(cache, ObjectIdGetDatum(address->objectId));
+
+		if (!HeapTupleIsValid(tuple))
+		{
+			return false;
+		}
+
+		ReleaseSysCache(tuple);
+
+		return true;
+	}
+	else
+	{
+		Relation	rel;
+		ScanKeyData skey[1];
+		SysScanDesc scan;
+		Snapshot	snapshot;
+
+		if (use_snapshot_self)
+			snapshot = SnapshotSelf;
+		else
+			snapshot = NULL;
+
+		rel = table_open(address->classId, AccessShareLock);
+
+		ScanKeyInit(&skey[0],
+					get_object_attnum_oid(address->classId),
+					BTEqualStrategyNumber, F_OIDEQ,
+					ObjectIdGetDatum(address->objectId));
+
+		scan = systable_beginscan(rel, get_object_oid_index(address->classId), true,
+								  snapshot, 1, skey);
+
+		/* we expect exactly one match */
+		tuple = systable_getnext(scan);
+		systable_endscan(scan);
+		table_close(rel, AccessShareLock);
+
+		return (HeapTupleIsValid(tuple));
+	}
+}
+
 /*
  * Return ObjectType for the given object type as given by
  * getObjectTypeDescription; if no valid ObjectType code exists, but it's a
diff --git a/src/backend/catalog/pg_depend.c b/src/backend/catalog/pg_depend.c
index 5366f7820c..f16af28429 100644
--- a/src/backend/catalog/pg_depend.c
+++ b/src/backend/catalog/pg_depend.c
@@ -108,6 +108,12 @@ recordMultipleDependencies(const ObjectAddress *depender,
 		if (isObjectPinned(referenced))
 			continue;
 
+		/*
+		 * Acquire a lock and check object still exists while recording the
+		 * dependency. XXX - Should we do so only for DEPENDENCY_NORMAL?
+		 */
+		depLockAndCheckObject(referenced);
+
 		if (slot_init_count < max_slots)
 		{
 			slot[slot_stored_count] = MakeSingleTupleTableSlot(RelationGetDescr(dependDesc),
@@ -506,6 +512,12 @@ changeDependencyFor(Oid classId, Oid objectId,
 		return 1;
 	}
 
+	/*
+	 * Acquire a lock and check object still exists while changing the
+	 * dependency.
+	 */
+	depLockAndCheckObject(&objAddr);
+
 	depRel = table_open(DependRelationId, RowExclusiveLock);
 
 	/* There should be existing dependency record(s), so search. */
diff --git a/src/backend/utils/errcodes.txt b/src/backend/utils/errcodes.txt
index 3250d539e1..60e8539fe3 100644
--- a/src/backend/utils/errcodes.txt
+++ b/src/backend/utils/errcodes.txt
@@ -271,6 +271,7 @@ Section: Class 28 - Invalid Authorization Specification
 Section: Class 2B - Dependent Privilege Descriptors Still Exist
 
 2B000    E    ERRCODE_DEPENDENT_PRIVILEGE_DESCRIPTORS_STILL_EXIST            dependent_privilege_descriptors_still_exist
+2BP02    E    ERRCODE_DEPENDENT_OBJECTS_DOES_NOT_EXIST                       dependent_objects_does_not_exist
 2BP01    E    ERRCODE_DEPENDENT_OBJECTS_STILL_EXIST                          dependent_objects_still_exist
 
 Section: Class 2D - Invalid Transaction Termination
diff --git a/src/include/catalog/dependency.h b/src/include/catalog/dependency.h
index 7eee66f810..5619b6f55e 100644
--- a/src/include/catalog/dependency.h
+++ b/src/include/catalog/dependency.h
@@ -101,6 +101,7 @@ typedef struct ObjectAddresses ObjectAddresses;
 /* in dependency.c */
 
 extern void AcquireDeletionLock(const ObjectAddress *object, int flags);
+extern void depLockAndCheckObject(const ObjectAddress *object);
 
 extern void ReleaseDeletionLock(const ObjectAddress *object);
 
diff --git a/src/include/catalog/objectaddress.h b/src/include/catalog/objectaddress.h
index 3a70d80e32..53ee9ad9e6 100644
--- a/src/include/catalog/objectaddress.h
+++ b/src/include/catalog/objectaddress.h
@@ -53,6 +53,7 @@ extern void check_object_ownership(Oid roleid,
 								   Node *object, Relation relation);
 
 extern Oid	get_object_namespace(const ObjectAddress *address);
+extern bool ObjectByIdExist(const ObjectAddress *address, bool use_snapshot_self);
 
 extern bool is_objectclass_supported(Oid class_id);
 extern const char *get_object_class_descr(Oid class_id);
diff --git a/src/test/isolation/expected/test_dependencies_locks.out b/src/test/isolation/expected/test_dependencies_locks.out
new file mode 100644
index 0000000000..9b645d7aa5
--- /dev/null
+++ b/src/test/isolation/expected/test_dependencies_locks.out
@@ -0,0 +1,129 @@
+Parsed test spec with 2 sessions
+
+starting permutation: s1_begin s1_create_function_in_schema s2_drop_schema s1_commit
+step s1_begin: BEGIN;
+step s1_create_function_in_schema: CREATE FUNCTION testschema.foo() RETURNS int AS 'select 1' LANGUAGE sql;
+step s2_drop_schema: DROP SCHEMA testschema; <waiting ...>
+step s1_commit: COMMIT;
+step s2_drop_schema: <... completed>
+ERROR:  cannot drop schema testschema because other objects depend on it
+
+starting permutation: s2_begin s2_drop_schema s1_create_function_in_schema s2_commit
+step s2_begin: BEGIN;
+step s2_drop_schema: DROP SCHEMA testschema;
+step s1_create_function_in_schema: CREATE FUNCTION testschema.foo() RETURNS int AS 'select 1' LANGUAGE sql; <waiting ...>
+step s2_commit: COMMIT;
+step s1_create_function_in_schema: <... completed>
+ERROR:  schema testschema does not exist
+
+starting permutation: s1_begin s1_alter_function_schema s2_drop_alterschema s1_commit
+step s1_begin: BEGIN;
+step s1_alter_function_schema: ALTER FUNCTION public.falter() SET SCHEMA alterschema;
+step s2_drop_alterschema: DROP SCHEMA alterschema; <waiting ...>
+step s1_commit: COMMIT;
+step s2_drop_alterschema: <... completed>
+ERROR:  cannot drop schema alterschema because other objects depend on it
+
+starting permutation: s2_begin s2_drop_alterschema s1_alter_function_schema s2_commit
+step s2_begin: BEGIN;
+step s2_drop_alterschema: DROP SCHEMA alterschema;
+step s1_alter_function_schema: ALTER FUNCTION public.falter() SET SCHEMA alterschema; <waiting ...>
+step s2_commit: COMMIT;
+step s1_alter_function_schema: <... completed>
+ERROR:  schema alterschema does not exist
+
+starting permutation: s1_begin s1_create_function_with_argtype s2_drop_foo_type s1_commit
+step s1_begin: BEGIN;
+step s1_create_function_with_argtype: CREATE FUNCTION fooargtype(num foo) RETURNS int AS 'select 1' LANGUAGE sql;
+step s2_drop_foo_type: DROP TYPE public.foo; <waiting ...>
+step s1_commit: COMMIT;
+step s2_drop_foo_type: <... completed>
+ERROR:  cannot drop type foo because other objects depend on it
+
+starting permutation: s2_begin s2_drop_foo_type s1_create_function_with_argtype s2_commit
+step s2_begin: BEGIN;
+step s2_drop_foo_type: DROP TYPE public.foo;
+step s1_create_function_with_argtype: CREATE FUNCTION fooargtype(num foo) RETURNS int AS 'select 1' LANGUAGE sql; <waiting ...>
+step s2_commit: COMMIT;
+step s1_create_function_with_argtype: <... completed>
+ERROR:  type foo does not exist
+
+starting permutation: s1_begin s1_create_function_with_rettype s2_drop_foo_rettype s1_commit
+step s1_begin: BEGIN;
+step s1_create_function_with_rettype: CREATE FUNCTION footrettype() RETURNS id LANGUAGE sql RETURN 1;
+step s2_drop_foo_rettype: DROP DOMAIN id; <waiting ...>
+step s1_commit: COMMIT;
+step s2_drop_foo_rettype: <... completed>
+ERROR:  cannot drop type id because other objects depend on it
+
+starting permutation: s2_begin s2_drop_foo_rettype s1_create_function_with_rettype s2_commit
+step s2_begin: BEGIN;
+step s2_drop_foo_rettype: DROP DOMAIN id;
+step s1_create_function_with_rettype: CREATE FUNCTION footrettype() RETURNS id LANGUAGE sql RETURN 1; <waiting ...>
+step s2_commit: COMMIT;
+step s1_create_function_with_rettype: <... completed>
+ERROR:  type id does not exist
+
+starting permutation: s1_begin s1_create_function_with_function s2_drop_function_f s1_commit
+step s1_begin: BEGIN;
+step s1_create_function_with_function: CREATE FUNCTION foofunc() RETURNS int LANGUAGE SQL RETURN f() + 1;
+step s2_drop_function_f: DROP FUNCTION f(); <waiting ...>
+step s1_commit: COMMIT;
+step s2_drop_function_f: <... completed>
+ERROR:  cannot drop function f() because other objects depend on it
+
+starting permutation: s2_begin s2_drop_function_f s1_create_function_with_function s2_commit
+step s2_begin: BEGIN;
+step s2_drop_function_f: DROP FUNCTION f();
+step s1_create_function_with_function: CREATE FUNCTION foofunc() RETURNS int LANGUAGE SQL RETURN f() + 1; <waiting ...>
+step s2_commit: COMMIT;
+step s1_create_function_with_function: <... completed>
+ERROR:  function f() does not exist
+
+starting permutation: s1_begin s1_create_domain_with_domain s2_drop_domain_id s1_commit
+step s1_begin: BEGIN;
+step s1_create_domain_with_domain: CREATE DOMAIN idid as id;
+step s2_drop_domain_id: DROP DOMAIN id; <waiting ...>
+step s1_commit: COMMIT;
+step s2_drop_domain_id: <... completed>
+ERROR:  cannot drop type id because other objects depend on it
+
+starting permutation: s2_begin s2_drop_domain_id s1_create_domain_with_domain s2_commit
+step s2_begin: BEGIN;
+step s2_drop_domain_id: DROP DOMAIN id;
+step s1_create_domain_with_domain: CREATE DOMAIN idid as id; <waiting ...>
+step s2_commit: COMMIT;
+step s1_create_domain_with_domain: <... completed>
+ERROR:  type id does not exist
+
+starting permutation: s1_begin s1_create_table_with_type s2_drop_footab_type s1_commit
+step s1_begin: BEGIN;
+step s1_create_table_with_type: CREATE TABLE tabtype(a footab);
+step s2_drop_footab_type: DROP TYPE public.footab; <waiting ...>
+step s1_commit: COMMIT;
+step s2_drop_footab_type: <... completed>
+ERROR:  cannot drop type footab because other objects depend on it
+
+starting permutation: s2_begin s2_drop_footab_type s1_create_table_with_type s2_commit
+step s2_begin: BEGIN;
+step s2_drop_footab_type: DROP TYPE public.footab;
+step s1_create_table_with_type: CREATE TABLE tabtype(a footab); <waiting ...>
+step s2_commit: COMMIT;
+step s1_create_table_with_type: <... completed>
+ERROR:  type footab does not exist
+
+starting permutation: s1_begin s1_create_server_with_fdw_wrapper s2_drop_fdw_wrapper s1_commit
+step s1_begin: BEGIN;
+step s1_create_server_with_fdw_wrapper: CREATE SERVER srv_fdw_wrapper FOREIGN DATA WRAPPER fdw_wrapper;
+step s2_drop_fdw_wrapper: DROP FOREIGN DATA WRAPPER fdw_wrapper RESTRICT; <waiting ...>
+step s1_commit: COMMIT;
+step s2_drop_fdw_wrapper: <... completed>
+ERROR:  cannot drop foreign-data wrapper fdw_wrapper because other objects depend on it
+
+starting permutation: s2_begin s2_drop_fdw_wrapper s1_create_server_with_fdw_wrapper s2_commit
+step s2_begin: BEGIN;
+step s2_drop_fdw_wrapper: DROP FOREIGN DATA WRAPPER fdw_wrapper RESTRICT;
+step s1_create_server_with_fdw_wrapper: CREATE SERVER srv_fdw_wrapper FOREIGN DATA WRAPPER fdw_wrapper; <waiting ...>
+step s2_commit: COMMIT;
+step s1_create_server_with_fdw_wrapper: <... completed>
+ERROR:  foreign-data wrapper fdw_wrapper does not exist
diff --git a/src/test/isolation/isolation_schedule b/src/test/isolation/isolation_schedule
index 0342eb39e4..1b67f0bffe 100644
--- a/src/test/isolation/isolation_schedule
+++ b/src/test/isolation/isolation_schedule
@@ -114,3 +114,4 @@ test: serializable-parallel-2
 test: serializable-parallel-3
 test: matview-write-skew
 test: lock-nowait
+test: test_dependencies_locks
diff --git a/src/test/isolation/specs/test_dependencies_locks.spec b/src/test/isolation/specs/test_dependencies_locks.spec
new file mode 100644
index 0000000000..5d04dfe9dc
--- /dev/null
+++ b/src/test/isolation/specs/test_dependencies_locks.spec
@@ -0,0 +1,89 @@
+setup
+{
+  CREATE SCHEMA testschema;
+  CREATE SCHEMA alterschema;
+  CREATE TYPE public.foo as enum ('one', 'two');
+  CREATE TYPE public.footab as enum ('three', 'four');
+  CREATE DOMAIN id AS int;
+  CREATE FUNCTION f() RETURNS int LANGUAGE SQL RETURN 1;
+  CREATE FUNCTION public.falter() RETURNS int LANGUAGE SQL RETURN 1;
+  CREATE FOREIGN DATA WRAPPER fdw_wrapper;
+}
+
+teardown
+{
+  DROP FUNCTION IF EXISTS testschema.foo();
+  DROP FUNCTION IF EXISTS fooargtype(num foo);
+  DROP FUNCTION IF EXISTS footrettype();
+  DROP FUNCTION IF EXISTS foofunc();
+  DROP FUNCTION IF EXISTS public.falter();
+  DROP FUNCTION IF EXISTS alterschema.falter();
+  DROP DOMAIN IF EXISTS idid;
+  DROP SERVER IF EXISTS srv_fdw_wrapper;
+  DROP TABLE IF EXISTS tabtype;
+  DROP SCHEMA IF EXISTS testschema;
+  DROP SCHEMA IF EXISTS alterschema;
+  DROP TYPE IF EXISTS public.foo;
+  DROP TYPE IF EXISTS public.footab;
+  DROP DOMAIN IF EXISTS id;
+  DROP FUNCTION IF EXISTS f();
+  DROP FOREIGN DATA WRAPPER IF EXISTS fdw_wrapper;
+}
+
+session "s1"
+
+step "s1_begin" { BEGIN; }
+step "s1_create_function_in_schema" { CREATE FUNCTION testschema.foo() RETURNS int AS 'select 1' LANGUAGE sql; }
+step "s1_create_function_with_argtype" { CREATE FUNCTION fooargtype(num foo) RETURNS int AS 'select 1' LANGUAGE sql; }
+step "s1_create_function_with_rettype" { CREATE FUNCTION footrettype() RETURNS id LANGUAGE sql RETURN 1; }
+step "s1_create_function_with_function" { CREATE FUNCTION foofunc() RETURNS int LANGUAGE SQL RETURN f() + 1; }
+step "s1_alter_function_schema" { ALTER FUNCTION public.falter() SET SCHEMA alterschema; }
+step "s1_create_domain_with_domain" { CREATE DOMAIN idid as id; }
+step "s1_create_table_with_type" { CREATE TABLE tabtype(a footab); }
+step "s1_create_server_with_fdw_wrapper" { CREATE SERVER srv_fdw_wrapper FOREIGN DATA WRAPPER fdw_wrapper; }
+step "s1_commit" { COMMIT; }
+
+session "s2"
+
+step "s2_begin" { BEGIN; }
+step "s2_drop_schema" { DROP SCHEMA testschema; }
+step "s2_drop_alterschema" { DROP SCHEMA alterschema; }
+step "s2_drop_foo_type" { DROP TYPE public.foo; }
+step "s2_drop_foo_rettype" { DROP DOMAIN id; }
+step "s2_drop_footab_type" { DROP TYPE public.footab; }
+step "s2_drop_function_f" { DROP FUNCTION f(); }
+step "s2_drop_domain_id" { DROP DOMAIN id; }
+step "s2_drop_fdw_wrapper" { DROP FOREIGN DATA WRAPPER fdw_wrapper RESTRICT; }
+step "s2_commit" { COMMIT; }
+
+# function - schema
+permutation "s1_begin" "s1_create_function_in_schema" "s2_drop_schema" "s1_commit"
+permutation "s2_begin" "s2_drop_schema" "s1_create_function_in_schema" "s2_commit"
+
+# alter function - schema
+permutation "s1_begin" "s1_alter_function_schema" "s2_drop_alterschema" "s1_commit"
+permutation "s2_begin" "s2_drop_alterschema" "s1_alter_function_schema" "s2_commit"
+
+# function - argtype
+permutation "s1_begin" "s1_create_function_with_argtype" "s2_drop_foo_type" "s1_commit"
+permutation "s2_begin" "s2_drop_foo_type" "s1_create_function_with_argtype" "s2_commit"
+
+# function - rettype
+permutation "s1_begin" "s1_create_function_with_rettype" "s2_drop_foo_rettype" "s1_commit"
+permutation "s2_begin" "s2_drop_foo_rettype" "s1_create_function_with_rettype" "s2_commit"
+
+# function - function
+permutation "s1_begin" "s1_create_function_with_function" "s2_drop_function_f" "s1_commit"
+permutation "s2_begin" "s2_drop_function_f" "s1_create_function_with_function" "s2_commit"
+
+# domain - domain
+permutation "s1_begin" "s1_create_domain_with_domain" "s2_drop_domain_id" "s1_commit"
+permutation "s2_begin" "s2_drop_domain_id" "s1_create_domain_with_domain" "s2_commit"
+
+# table - type
+permutation "s1_begin" "s1_create_table_with_type" "s2_drop_footab_type" "s1_commit"
+permutation "s2_begin" "s2_drop_footab_type" "s1_create_table_with_type" "s2_commit"
+
+# server - foreign data wrapper
+permutation "s1_begin" "s1_create_server_with_fdw_wrapper" "s2_drop_fdw_wrapper" "s1_commit"
+permutation "s2_begin" "s2_drop_fdw_wrapper" "s1_create_server_with_fdw_wrapper" "s2_commit"
-- 
2.34.1


--Z+BFS5G1lhqjvsq4--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread

* [PATCH] Teach REPACK to upgrade its lock safely.
@ 2026-04-14 09:59 Antonin Houska <ah@cybertec.at>
  0 siblings, 0 replies; 399+ messages in thread

From: Antonin Houska @ 2026-04-14 09:59 UTC (permalink / raw)

REPACK (CONCURRENTLY) needs to upgrade its ShareUpdateExclusiveLock (SUEL) to
AccessExclusiveLock at the end of table processing. If another session, which
already has a non-conflicting lock on the table, tried to get SUEL too, it
would end up in a deadlock with REPACK. Such situation does not hurt as long
as the deadlock detector choses to terminate the other session, but it's
possible that it terminates REPACK. A lot of REPACK's work would be wasted
this way.

This patch checks for such situation, and if the other session tries to get
the SUEL, it receives a deadlock report before it actually starts waiting.

Please note that the other session can safely run VACUUM w/o receiving the
deadlock report. The point is that VACUUM cannot run in a transaction block,
so the session cannot have any other lock when trying to get SUEL on the
table. In that case, REPACK gets its AccessExclusiveLock first, even if VACUUM
requested SUEL earlier. VACUUM then just needs wait for REPACK to finish, and
then it can start.
---
 src/backend/commands/repack.c                 | 25 +++++++-
 src/backend/storage/lmgr/deadlock.c           | 10 ++--
 src/backend/storage/lmgr/lock.c               | 30 ++++++++++
 src/backend/storage/lmgr/proc.c               | 55 ++++++++++++++++-
 src/include/storage/lock.h                    |  5 +-
 src/include/storage/proc.h                    |  6 +-
 .../injection_points/expected/repack.out      | 59 ++++++++++++++++++-
 .../injection_points/specs/repack.spec        | 29 +++++++++
 8 files changed, 209 insertions(+), 10 deletions(-)

diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c
index 58e3867246f..e9193067666 100644
--- a/src/backend/commands/repack.c
+++ b/src/backend/commands/repack.c
@@ -285,6 +285,18 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel)
 		 * to understand and we don't lose any functionality.
 		 */
 		PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)");
+
+		/*
+		 * Also set the PROC_IN_CONCURRENT_REPACK flag.  This makes the lock
+		 * manager cause anyone that would conflict with us to error out.
+		 * It's important to set this flag ahead of actually locking the
+		 * relation; it won't of course affect anyone until we do have a lock
+		 * that others can conflict with.
+		 */
+		LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+		MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK;
+		ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+		LWLockRelease(ProcArrayLock);
 	}
 
 	/*
@@ -3086,7 +3098,18 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap,
 		LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock);
 
 	/*
-	 * Tuples and pages of the old heap will be gone, but the heap will stay.
+	 * Now that we have all access-exclusive locks on all relations, we no
+	 * longer want other processes to error out when trying to acquire a
+	 * conflicting lock.  Therefore, unset our flag.
+	 */
+	LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE);
+	MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK;
+	ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags;
+	LWLockRelease(ProcArrayLock);
+
+	/*
+	 * Tuples and pages of the old heap will be gone, but the heap itself will
+	 * stay.
 	 */
 	TransferPredicateLocksToHeapRelation(OldHeap);
 	foreach_ptr(RelationData, index, indexrels)
diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c
index b8962d875b6..3772a7aa3df 100644
--- a/src/backend/storage/lmgr/deadlock.c
+++ b/src/backend/storage/lmgr/deadlock.c
@@ -1146,8 +1146,10 @@ DeadLockReport(void)
 void
 RememberSimpleDeadLock(PGPROC *proc1,
 					   LOCKMODE lockmode,
-					   LOCK *lock,
-					   PGPROC *proc2)
+					   LOCK *lock, /* XXX Only lockmode? */
+					   PGPROC *proc2,
+					   LOCKMODE lockmode2,
+					   LOCKTAG *locktag2)
 {
 	DEADLOCK_INFO *info = &deadlockDetails[0];
 
@@ -1155,8 +1157,8 @@ RememberSimpleDeadLock(PGPROC *proc1,
 	info->lockmode = lockmode;
 	info->pid = proc1->pid;
 	info++;
-	info->locktag = proc2->waitLock->tag;
-	info->lockmode = proc2->waitLockMode;
+	info->locktag = *locktag2;
+	info->lockmode = lockmode2;
 	info->pid = proc2->pid;
 	nDeadlockDetails = 2;
 }
diff --git a/src/backend/storage/lmgr/lock.c b/src/backend/storage/lmgr/lock.c
index c221fe96889..44aa2ca26a7 100644
--- a/src/backend/storage/lmgr/lock.c
+++ b/src/backend/storage/lmgr/lock.c
@@ -674,6 +674,36 @@ LockHeldByMe(const LOCKTAG *locktag,
 	return false;
 }
 
+/*
+ * HaveRelLock -- test whether the current backend has a fast path lock on
+ *		the relation in any mode.
+ */
+bool
+HaveFastPathLock(Oid relid)
+{
+	uint32		group = FAST_PATH_REL_GROUP(relid);
+	bool		result = false;
+
+	LWLockAcquire(&MyProc->fpInfoLock, LW_SHARED);
+	for (int i = 0; i < FP_LOCK_SLOTS_PER_GROUP; i++)
+	{
+		uint32		f = FAST_PATH_SLOT(group, i);
+		uint32		lockmask;
+
+		if (relid != MyProc->fpRelId[f])
+			continue;
+		lockmask = FAST_PATH_GET_BITS(MyProc, f);
+		if (lockmask)
+		{
+			result = true;
+			break;
+		}
+	}
+	LWLockRelease(&MyProc->fpInfoLock);
+
+	return result;
+}
+
 #ifdef USE_ASSERT_CHECKING
 /*
  * GetLockMethodLocalHash -- return the hash of local locks, for modules that
diff --git a/src/backend/storage/lmgr/proc.c b/src/backend/storage/lmgr/proc.c
index 1ac25068d62..3d47b1284a3 100644
--- a/src/backend/storage/lmgr/proc.c
+++ b/src/backend/storage/lmgr/proc.c
@@ -1156,6 +1156,7 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 	LOCKMASK	myHeldLocks;
 	bool		early_deadlock = false;
 	PGPROC	   *leader = MyProc->lockGroupLeader;
+	bool		check_repack = false;
 
 	Assert(LWLockHeldByMeInMode(partitionLock, LW_EXCLUSIVE));
 
@@ -1193,6 +1194,56 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 		}
 	}
 
+	/*
+	 * If I am already holding this lock (in any mode) and trying to get
+	 * ShareUpdateExclusiveLock (or higher), check if REPACK (CONCURRENTLY) is
+	 * already holding ShareUpdateExclusiveLock. The problem is if that by
+	 * trying to upgrade the lock to AccessExclusiveLock it would get into a
+	 * deadlock with us.
+	 */
+	if (lock->tag.locktag_type == LOCKTAG_RELATION &&
+		lock->tag.locktag_field1 == MyDatabaseId &&
+		lockmode >= ShareUpdateExclusiveLock)
+	{
+		if (myHeldLocks > 0 ||
+			HaveFastPathLock(lock->tag.locktag_field2))
+			check_repack = true;
+	}
+	if (check_repack)
+	{
+		dlist_iter	iter;
+
+		dlist_foreach(iter, &lock->procLocks)
+		{
+			PROCLOCK   *otherproclock;
+			PGPROC		*otherproc;
+
+			otherproclock = dlist_container(PROCLOCK, lockLink, iter.cur);
+			otherproc = otherproclock->tag.myProc;
+			if (otherproc != MyProc &&
+				otherproc->statusFlags & PROC_IN_CONCURRENT_REPACK)
+			{
+				LOCKMASK	repackmask = otherproclock->holdMask;
+
+				/*
+				 * Is REPACK already holding ShareUpdateExclusiveLock?
+				 */
+				if ((repackmask & LOCKBIT_ON(ShareUpdateExclusiveLock)) != 0)
+				{
+					/*
+					 * There should be no more than one REPACK working on
+					 * particular table, so let's error out.
+					 */
+					RememberSimpleDeadLock(MyProc, lockmode, lock,
+										   otherproc,
+										   ShareUpdateExclusiveLock,
+										   &otherproclock->tag.myLock->tag);
+					return PROC_WAIT_STATUS_ERROR;
+				}
+			}
+		}
+	}
+
 	/*
 	 * Determine where to add myself in the wait queue.
 	 *
@@ -1240,7 +1291,9 @@ JoinWaitQueue(LOCALLOCK *locallock, LockMethod lockMethodTable, bool dontWait)
 					 * a flag to check below, and break out of loop.  Also,
 					 * record deadlock info for later message.
 					 */
-					RememberSimpleDeadLock(MyProc, lockmode, lock, proc);
+					RememberSimpleDeadLock(MyProc, lockmode, lock, proc,
+										   proc->waitLockMode,
+										   &proc->waitLock->tag);
 					early_deadlock = true;
 					break;
 				}
diff --git a/src/include/storage/lock.h b/src/include/storage/lock.h
index ee3cb1dc203..85804d61ce9 100644
--- a/src/include/storage/lock.h
+++ b/src/include/storage/lock.h
@@ -401,6 +401,7 @@ extern void LockReleaseCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern void LockReassignCurrentOwner(LOCALLOCK **locallocks, int nlocks);
 extern bool LockHeldByMe(const LOCKTAG *locktag,
 						 LOCKMODE lockmode, bool orstronger);
+extern bool HaveFastPathLock(Oid relid);
 #ifdef USE_ASSERT_CHECKING
 extern HTAB *GetLockMethodLocalHash(void);
 #endif
@@ -440,7 +441,9 @@ pg_noreturn extern void DeadLockReport(void);
 extern void RememberSimpleDeadLock(PGPROC *proc1,
 								   LOCKMODE lockmode,
 								   LOCK *lock,
-								   PGPROC *proc2);
+								   PGPROC *proc2,
+								   LOCKMODE lockmode2,
+								   LOCKTAG *locktag2);
 extern void InitDeadLockChecking(void);
 
 extern int	LockWaiterCount(const LOCKTAG *locktag);
diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h
index 3e1d1fad5f9..76c6bb44251 100644
--- a/src/include/storage/proc.h
+++ b/src/include/storage/proc.h
@@ -70,10 +70,12 @@ struct XidCache
 #define		PROC_AFFECTS_ALL_HORIZONS	0x20	/* this proc's xmin must be
 												 * included in vacuum horizons
 												 * in all databases */
+#define		PROC_IN_CONCURRENT_REPACK	0x40	/* REPACK (CONCURRENTLY) */
 
-/* flags reset at EOXact */
+/* flags reset at EOXact.  A bit of a misnomer ... */
 #define		PROC_VACUUM_STATE_MASK \
-	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND)
+	(PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \
+	 PROC_IN_CONCURRENT_REPACK)
 
 /*
  * Xmin-related flags. Make sure any flags that affect how the process' Xmin
diff --git a/src/test/modules/injection_points/expected/repack.out b/src/test/modules/injection_points/expected/repack.out
index b575e9052ee..960044776f4 100644
--- a/src/test/modules/injection_points/expected/repack.out
+++ b/src/test/modules/injection_points/expected/repack.out
@@ -1,4 +1,4 @@
-Parsed test spec with 2 sessions
+Parsed test spec with 3 sessions
 
 starting permutation: wait_before_lock change_existing change_new change_subxact1 change_subxact2 check2 wakeup_before_lock check1
 injection_points_attach
@@ -111,3 +111,60 @@ injection_points_detach
                        
 (1 row)
 
+
+starting permutation: check1_relnode_only wait_before_lock lock3 wakeup_before_lock check1_relnode_only
+injection_points_attach
+-----------------------
+                       
+(1 row)
+
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    1
+(1 row)
+
+step wait_before_lock: 
+	REPACK (CONCURRENTLY) repack_test USING INDEX repack_test_pkey;
+ <waiting ...>
+step lock3: 
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+
+i|j
+-+-
+1|1
+(1 row)
+
+ERROR:  deadlock detected
+step wakeup_before_lock: 
+	SELECT injection_points_wakeup('repack-concurrently-before-lock');
+
+injection_points_wakeup
+-----------------------
+                       
+(1 row)
+
+step wait_before_lock: <... completed>
+step check1_relnode_only: 
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+
+count
+-----
+    2
+(1 row)
+
+injection_points_detach
+-----------------------
+                       
+(1 row)
+
diff --git a/src/test/modules/injection_points/specs/repack.spec b/src/test/modules/injection_points/specs/repack.spec
index d727a9b056b..5ffcb558718 100644
--- a/src/test/modules/injection_points/specs/repack.spec
+++ b/src/test/modules/injection_points/specs/repack.spec
@@ -55,6 +55,13 @@ step check1
 	FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j)
 	WHERE d1.i ISNULL OR d2.i ISNULL;
 }
+step check1_relnode_only
+{
+	INSERT INTO relfilenodes(node)
+	SELECT relfilenode FROM pg_class WHERE relname='repack_test';
+
+	SELECT count(DISTINCT node) FROM relfilenodes;
+}
 teardown
 {
 	SELECT injection_points_detach('repack-concurrently-before-lock');
@@ -129,6 +136,20 @@ step wakeup_before_lock
 	SELECT injection_points_wakeup('repack-concurrently-before-lock');
 }
 
+session s3
+# Try to acquire SharedUpdateExclusiveLock on a table while REPACK is already
+# holding it and before it tries to upgrade it to AccessExclusiveLock. Since
+# this session is already holding another lock no the table, REPACK cannot
+# resolve the problem by getting ahead in the wait queue. The solution is that
+# the LOCK TABLE triggers a deadlock report. The deadlock does not actually
+# take place, but it would if this session started to wait.
+step lock3
+{
+	BEGIN;
+	SELECT * FROM repack_test ORDER BY i LIMIT 1;
+	LOCK TABLE repack_test IN SHARE UPDATE EXCLUSIVE MODE;
+}
+
 # Test if data changes introduced while one session is performing REPACK
 # CONCURRENTLY find their way into the table.
 permutation
@@ -140,3 +161,11 @@ permutation
 	check2
 	wakeup_before_lock
 	check1
+
+# See 'lock3' above. We check relnodes to make sure that REPACK finished.
+permutation
+	check1_relnode_only
+	wait_before_lock
+	lock3
+	wakeup_before_lock
+	check1_relnode_only
-- 
2.47.3


--=-=-=--





^ permalink  raw  reply  [nested|flat] 399+ messages in thread


end of thread, other threads:[~2026-04-14 09:59 UTC | newest]

Thread overview: 399+ messages (download: mbox mbox.gz follow: Atom feed)
-- links below jump to the message on this page --
2024-03-29 15:43 [PATCH v8] Avoid orphaned objects dependencies Bertrand Drouvot <bertranddrouvot.pg@gmail.com>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>
2026-04-14 09:59 [PATCH] Teach REPACK to upgrade its lock safely. Antonin Houska <ah@cybertec.at>

This inbox is served by agora; see mirroring instructions
for how to clone and mirror all data and code used for this inbox